[Core] Scope PCP-DP validation to GPU manager - #54523
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
acd364b to
29ce06a
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: acd364b85a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
/ci run |
|
✅ Triggered Buildkite CI #86328 for commit |
|
✅ @pisceskkk, CI is now available for this PR.
|
|
/ci cancel |
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
|
/ci retry |
|
✅ Queued 2 failed job(s) for retry in Buildkite CI #86747. |
|
/amd-ci retry |
|
✅ Queued 1 failed job(s) for retry in Buildkite AMD CI #12538. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Team Run ID: 📒 Files selected for processing (3)
💤 Files with no reviewable changes (1)
🚧 Files skipped from review as they are similar to previous changes (2)
Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review. 📝 SummarySummary by CodeRabbit
WalkthroughThe parallel configuration validator now permits PCP with data parallelism. CUDA and ROCm platform validators reject this combination with platform-specific ChangesPCP and data parallelism validation
Estimated code review effort: 2 (Simple) | ~10 minutes Merge Risk: ⚪ Minimal · up to PCP and data-parallelism validation is now enforced by CUDA and ROCm platform checks while allowing other backends to define their own support. No current merge-blocking risk is identified. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
/ci retry |
|
✅ The previous CI build is still running: https://buildkite.com/vllm/ci/builds/86747 |
|
/amd-ci retry |
|
✅ Queued 3 failed job(s) for retry in Buildkite AMD CI #12631. |
|
Note GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer. |
Rely on vllm-project/vllm#54523 to scope PCP+DP rejection to CUDA and ROCm. Remove the Ascend validator wrapper, cached schema rebuilding, dedicated tests and patch documentation. Keep dummy execution handling. Validation: targeted Ruff checks, syntax checks and git diff --check. Runtime compatibility with vLLM 7c2f1ff remains pending. Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
### What this PR does / why we need it? Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream PCP dispatch sizing and input partitioning paths. When a DP replica becomes idle, its dummy batch bypasses PCP partitioning. Saved PCP attention state may be absent or belong to the preceding real batch. Runtime dummy inputs use runner buffers, while decode graphs capture persistent PCP buffers. - Build dummy PCP attention contexts from the current batch, block tables and rank-local slot mappings, including the DSA metadata path. - Refresh persistent PCP buffers before dummy execution. Keep the implementation in `AscendPCPManager`; `NPUModelRunner.prepare_dummy_attn` delegates or takes the existing non-PCP path. - Synchronize replicated speculative drafts independently of the target's PCP-local DP state, with every DP replica participating, including idle and decode replicas. - Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`. Upstream validation dependency: vllm-project/vllm#54523, merged as [7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff), moves the PCP+DP restriction from common `ParallelConfig` validation to CUDA/ROCm platform checks. This PR relies on that change and contains no Ascend validator bypass or Pydantic schema rebuilding. Related upstream PR: vllm-project/vllm#53867 ([Feature][PCP] Support decode-only FULL CUDA graphs). Once the paired vLLM includes #53867, adapt its PCP `prepare_inputs_to_capture` entry point to create `AscendInputBatch` directly in persistent PCP buffers. After capture and real-to-idle replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy attention-context handling and replicated-draft DP synchronization remain necessary. ### Does this PR introduce _any_ user-facing change? Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI options or environment variables are introduced. Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp` forwarding and applies fresh synchronization for replicated drafts on the main2main `dp_sync` interface. **Compatibility and remaining validation:** - The previously tested baseline was Ascend `b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM `e6bfe03ad73a3330cb427885aa90d97a12e1c704`. - Removing the validator bypass requires vLLM #54523. Ascend main currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which predates that change and still rejects PCP+DP. The local vLLM checkout remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. This PR does not change the main2main pin or reintroduce the bypass. - Compatibility work against that newer revision is still pending, including the added `prepare_dummy_attn` argument and removal of `InputBatch.max_seq_len_np`. The current branch is not yet validated to run against that revision. - Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions during `FULL_DECODE_ONLY` replay, speculative-decoding numerical correctness, and EP/TP combinations require model-level validation on the final paired sources. ### How was this patch tested? The latest rebase onto `ad86348b0` resolves an append-location conflict in `test_pcp_manager_v2.py`, retaining both the main DCP padding regression and this PR's tests. AST comparison confirms that all main tests and the final PR tests are preserved, and the other changed files match the automatic merge. Ruff lint/format, syntax parsing, and `git diff --check` passed. Runtime tests were not rerun for this conflict-only update. For the preceding rebase onto `f8481287a`, **6 targeted candidate-method cases passed** in an isolated process in the Ascend container: five main2main draft-sync cases and one mocked v0.28.0 argument-routing case. The exact candidate `propose` method was loaded into the installed runtime classes; native dependencies and the rest of the runtime remained at their installed versions. This is method-level regression evidence, not full rebased-source or dual-version runtime validation. All eight changed Python files passed Ruff lint/format checks, syntax parsing, and `git diff --check`. Full unit-suite and model validation on the final paired sources remain pending. The runtime adaptation suite previously passed on the earlier baseline in an isolated Ascend-container test directory, using CPU tensors and mocks for device kernels: ```bash python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py ``` - **75 passed** (14 existing `torch.jit` deprecation warnings). This is prior-baseline evidence, not a test result for vLLM `7c2f1ff`. - Covers missing/stale dummy PCP state, rank-local mappings, persistent buffer contents and storage, DSA metadata forwarding, non-PCP fallback, graph restrictions, and replicated draft DP synchronization. - The removed validator bypass's dedicated tests are also removed; the existing unrelated platform test is retained. - For the removal, targeted Ruff lint/format checks, Python syntax checks and `git diff --check` passed. No runtime tests were rerun against the upgraded vLLM. The local `format.sh ci` entry point remains unavailable because its shell lacks `pre-commit`. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
### What this PR does / why we need it? Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream PCP dispatch sizing and input partitioning paths. When a DP replica becomes idle, its dummy batch bypasses PCP partitioning. Saved PCP attention state may be absent or belong to the preceding real batch. Runtime dummy inputs use runner buffers, while decode graphs capture persistent PCP buffers. - Build dummy PCP attention contexts from the current batch, block tables and rank-local slot mappings, including the DSA metadata path. - Refresh persistent PCP buffers before dummy execution. Keep the implementation in `AscendPCPManager`; `NPUModelRunner.prepare_dummy_attn` delegates or takes the existing non-PCP path. - Synchronize replicated speculative drafts independently of the target's PCP-local DP state, with every DP replica participating, including idle and decode replicas. - Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`. Upstream validation dependency: vllm-project/vllm#54523, merged as [7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff), moves the PCP+DP restriction from common `ParallelConfig` validation to CUDA/ROCm platform checks. This PR relies on that change and contains no Ascend validator bypass or Pydantic schema rebuilding. Related upstream PR: vllm-project/vllm#53867 ([Feature][PCP] Support decode-only FULL CUDA graphs). Once the paired vLLM includes #53867, adapt its PCP `prepare_inputs_to_capture` entry point to create `AscendInputBatch` directly in persistent PCP buffers. After capture and real-to-idle replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy attention-context handling and replicated-draft DP synchronization remain necessary. ### Does this PR introduce _any_ user-facing change? Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI options or environment variables are introduced. Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp` forwarding and applies fresh synchronization for replicated drafts on the main2main `dp_sync` interface. **Compatibility and remaining validation:** - The previously tested baseline was Ascend `b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM `e6bfe03ad73a3330cb427885aa90d97a12e1c704`. - Removing the validator bypass requires vLLM #54523. Ascend main currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which predates that change and still rejects PCP+DP. The local vLLM checkout remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. This PR does not change the main2main pin or reintroduce the bypass. - Compatibility work against that newer revision is still pending, including the added `prepare_dummy_attn` argument and removal of `InputBatch.max_seq_len_np`. The current branch is not yet validated to run against that revision. - Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions during `FULL_DECODE_ONLY` replay, speculative-decoding numerical correctness, and EP/TP combinations require model-level validation on the final paired sources. ### How was this patch tested? The latest rebase onto `ad86348b0` resolves an append-location conflict in `test_pcp_manager_v2.py`, retaining both the main DCP padding regression and this PR's tests. AST comparison confirms that all main tests and the final PR tests are preserved, and the other changed files match the automatic merge. Ruff lint/format, syntax parsing, and `git diff --check` passed. Runtime tests were not rerun for this conflict-only update. For the preceding rebase onto `f8481287a`, **6 targeted candidate-method cases passed** in an isolated process in the Ascend container: five main2main draft-sync cases and one mocked v0.28.0 argument-routing case. The exact candidate `propose` method was loaded into the installed runtime classes; native dependencies and the rest of the runtime remained at their installed versions. This is method-level regression evidence, not full rebased-source or dual-version runtime validation. All eight changed Python files passed Ruff lint/format checks, syntax parsing, and `git diff --check`. Full unit-suite and model validation on the final paired sources remain pending. The runtime adaptation suite previously passed on the earlier baseline in an isolated Ascend-container test directory, using CPU tensors and mocks for device kernels: ```bash python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py ``` - **75 passed** (14 existing `torch.jit` deprecation warnings). This is prior-baseline evidence, not a test result for vLLM `7c2f1ff`. - Covers missing/stale dummy PCP state, rank-local mappings, persistent buffer contents and storage, DSA metadata forwarding, non-PCP fallback, graph restrictions, and replicated draft DP synchronization. - The removed validator bypass's dedicated tests are also removed; the existing unrelated platform test is retained. - For the removal, targeted Ruff lint/format checks, Python syntax checks and `git diff --check` passed. No runtime tests were rerun against the upgraded vLLM. The local `format.sh ci` entry point remains unavailable because its shell lacks `pre-commit`. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com> Signed-off-by: Jyotirmoy Roy <jyotirmoyroy649@gmail.com>
### What this PR does / why we need it? Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream PCP dispatch sizing and input partitioning paths. When a DP replica becomes idle, its dummy batch bypasses PCP partitioning. Saved PCP attention state may be absent or belong to the preceding real batch. Runtime dummy inputs use runner buffers, while decode graphs capture persistent PCP buffers. - Build dummy PCP attention contexts from the current batch, block tables and rank-local slot mappings, including the DSA metadata path. - Refresh persistent PCP buffers before dummy execution. Keep the implementation in `AscendPCPManager`; `NPUModelRunner.prepare_dummy_attn` delegates or takes the existing non-PCP path. - Synchronize replicated speculative drafts independently of the target's PCP-local DP state, with every DP replica participating, including idle and decode replicas. - Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`. Upstream validation dependency: vllm-project/vllm#54523, merged as [7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff), moves the PCP+DP restriction from common `ParallelConfig` validation to CUDA/ROCm platform checks. This PR relies on that change and contains no Ascend validator bypass or Pydantic schema rebuilding. Related upstream PR: vllm-project/vllm#53867 ([Feature][PCP] Support decode-only FULL CUDA graphs). Once the paired vLLM includes #53867, adapt its PCP `prepare_inputs_to_capture` entry point to create `AscendInputBatch` directly in persistent PCP buffers. After capture and real-to-idle replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy attention-context handling and replicated-draft DP synchronization remain necessary. ### Does this PR introduce _any_ user-facing change? Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI options or environment variables are introduced. Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp` forwarding and applies fresh synchronization for replicated drafts on the main2main `dp_sync` interface. **Compatibility and remaining validation:** - The previously tested baseline was Ascend `b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM `e6bfe03ad73a3330cb427885aa90d97a12e1c704`. - Removing the validator bypass requires vLLM #54523. Ascend main currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which predates that change and still rejects PCP+DP. The local vLLM checkout remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. This PR does not change the main2main pin or reintroduce the bypass. - Compatibility work against that newer revision is still pending, including the added `prepare_dummy_attn` argument and removal of `InputBatch.max_seq_len_np`. The current branch is not yet validated to run against that revision. - Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions during `FULL_DECODE_ONLY` replay, speculative-decoding numerical correctness, and EP/TP combinations require model-level validation on the final paired sources. ### How was this patch tested? The latest rebase onto `ad86348b0` resolves an append-location conflict in `test_pcp_manager_v2.py`, retaining both the main DCP padding regression and this PR's tests. AST comparison confirms that all main tests and the final PR tests are preserved, and the other changed files match the automatic merge. Ruff lint/format, syntax parsing, and `git diff --check` passed. Runtime tests were not rerun for this conflict-only update. For the preceding rebase onto `f8481287a`, **6 targeted candidate-method cases passed** in an isolated process in the Ascend container: five main2main draft-sync cases and one mocked v0.28.0 argument-routing case. The exact candidate `propose` method was loaded into the installed runtime classes; native dependencies and the rest of the runtime remained at their installed versions. This is method-level regression evidence, not full rebased-source or dual-version runtime validation. All eight changed Python files passed Ruff lint/format checks, syntax parsing, and `git diff --check`. Full unit-suite and model validation on the final paired sources remain pending. The runtime adaptation suite previously passed on the earlier baseline in an isolated Ascend-container test directory, using CPU tensors and mocks for device kernels: ```bash python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py ``` - **75 passed** (14 existing `torch.jit` deprecation warnings). This is prior-baseline evidence, not a test result for vLLM `7c2f1ff`. - Covers missing/stale dummy PCP state, rank-local mappings, persistent buffer contents and storage, DSA metadata forwarding, non-PCP fallback, graph restrictions, and replicated draft DP synchronization. - The removed validator bypass's dedicated tests are also removed; the existing unrelated platform test is retained. - For the removal, targeted Ruff lint/format checks, Python syntax checks and `git diff --check` passed. No runtime tests were rerun against the upgraded vLLM. The local `format.sh ci` entry point remains unavailable because its shell lacks `pre-commit`. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
### What this PR does / why we need it? Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream PCP dispatch sizing and input partitioning paths. When a DP replica becomes idle, its dummy batch bypasses PCP partitioning. Saved PCP attention state may be absent or belong to the preceding real batch. Runtime dummy inputs use runner buffers, while decode graphs capture persistent PCP buffers. - Build dummy PCP attention contexts from the current batch, block tables and rank-local slot mappings, including the DSA metadata path. - Refresh persistent PCP buffers before dummy execution. Keep the implementation in `AscendPCPManager`; `NPUModelRunner.prepare_dummy_attn` delegates or takes the existing non-PCP path. - Synchronize replicated speculative drafts independently of the target's PCP-local DP state, with every DP replica participating, including idle and decode replicas. - Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`. Upstream validation dependency: vllm-project/vllm#54523, merged as [7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff), moves the PCP+DP restriction from common `ParallelConfig` validation to CUDA/ROCm platform checks. This PR relies on that change and contains no Ascend validator bypass or Pydantic schema rebuilding. Related upstream PR: vllm-project/vllm#53867 ([Feature][PCP] Support decode-only FULL CUDA graphs). Once the paired vLLM includes #53867, adapt its PCP `prepare_inputs_to_capture` entry point to create `AscendInputBatch` directly in persistent PCP buffers. After capture and real-to-idle replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy attention-context handling and replicated-draft DP synchronization remain necessary. ### Does this PR introduce _any_ user-facing change? Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI options or environment variables are introduced. Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp` forwarding and applies fresh synchronization for replicated drafts on the main2main `dp_sync` interface. **Compatibility and remaining validation:** - The previously tested baseline was Ascend `b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM `e6bfe03ad73a3330cb427885aa90d97a12e1c704`. - Removing the validator bypass requires vLLM #54523. Ascend main currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which predates that change and still rejects PCP+DP. The local vLLM checkout remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. This PR does not change the main2main pin or reintroduce the bypass. - Compatibility work against that newer revision is still pending, including the added `prepare_dummy_attn` argument and removal of `InputBatch.max_seq_len_np`. The current branch is not yet validated to run against that revision. - Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions during `FULL_DECODE_ONLY` replay, speculative-decoding numerical correctness, and EP/TP combinations require model-level validation on the final paired sources. ### How was this patch tested? The latest rebase onto `ad86348b0` resolves an append-location conflict in `test_pcp_manager_v2.py`, retaining both the main DCP padding regression and this PR's tests. AST comparison confirms that all main tests and the final PR tests are preserved, and the other changed files match the automatic merge. Ruff lint/format, syntax parsing, and `git diff --check` passed. Runtime tests were not rerun for this conflict-only update. For the preceding rebase onto `f8481287a`, **6 targeted candidate-method cases passed** in an isolated process in the Ascend container: five main2main draft-sync cases and one mocked v0.28.0 argument-routing case. The exact candidate `propose` method was loaded into the installed runtime classes; native dependencies and the rest of the runtime remained at their installed versions. This is method-level regression evidence, not full rebased-source or dual-version runtime validation. All eight changed Python files passed Ruff lint/format checks, syntax parsing, and `git diff --check`. Full unit-suite and model validation on the final paired sources remain pending. The runtime adaptation suite previously passed on the earlier baseline in an isolated Ascend-container test directory, using CPU tensors and mocks for device kernels: ```bash python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py ``` - **75 passed** (14 existing `torch.jit` deprecation warnings). This is prior-baseline evidence, not a test result for vLLM `7c2f1ff`. - Covers missing/stale dummy PCP state, rank-local mappings, persistent buffer contents and storage, DSA metadata forwarding, non-PCP fallback, graph restrictions, and replicated draft DP synchronization. - The removed validator bypass's dedicated tests are also removed; the existing unrelated platform test is retained. - For the removal, targeted Ruff lint/format checks, Python syntax checks and `git diff --check` passed. No runtime tests were rerun against the upgraded vLLM. The local `format.sh ci` entry point remains unavailable because its shell lacks `pre-commit`. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wzx0726 <zhexuanwu12@gmail.com> Signed-off-by: tianming2009 <13246728590@163.com>
vllm-project#54523 moved the ParallelConfig check into CudaPlatform/RocmPlatform check_and_update_config after this series branched; the series removes the ParallelConfig copy, so remove the platform copies too. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
### What this PR does / why we need it? Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream PCP dispatch sizing and input partitioning paths. When a DP replica becomes idle, its dummy batch bypasses PCP partitioning. Saved PCP attention state may be absent or belong to the preceding real batch. Runtime dummy inputs use runner buffers, while decode graphs capture persistent PCP buffers. - Build dummy PCP attention contexts from the current batch, block tables and rank-local slot mappings, including the DSA metadata path. - Refresh persistent PCP buffers before dummy execution. Keep the implementation in `AscendPCPManager`; `NPUModelRunner.prepare_dummy_attn` delegates or takes the existing non-PCP path. - Synchronize replicated speculative drafts independently of the target's PCP-local DP state, with every DP replica participating, including idle and decode replicas. - Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`. Upstream validation dependency: vllm-project/vllm#54523, merged as [7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff), moves the PCP+DP restriction from common `ParallelConfig` validation to CUDA/ROCm platform checks. This PR relies on that change and contains no Ascend validator bypass or Pydantic schema rebuilding. Related upstream PR: vllm-project/vllm#53867 ([Feature][PCP] Support decode-only FULL CUDA graphs). Once the paired vLLM includes #53867, adapt its PCP `prepare_inputs_to_capture` entry point to create `AscendInputBatch` directly in persistent PCP buffers. After capture and real-to-idle replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy attention-context handling and replicated-draft DP synchronization remain necessary. ### Does this PR introduce _any_ user-facing change? Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI options or environment variables are introduced. Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp` forwarding and applies fresh synchronization for replicated drafts on the main2main `dp_sync` interface. **Compatibility and remaining validation:** - The previously tested baseline was Ascend `b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM `e6bfe03ad73a3330cb427885aa90d97a12e1c704`. - Removing the validator bypass requires vLLM #54523. Ascend main currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which predates that change and still rejects PCP+DP. The local vLLM checkout remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. This PR does not change the main2main pin or reintroduce the bypass. - Compatibility work against that newer revision is still pending, including the added `prepare_dummy_attn` argument and removal of `InputBatch.max_seq_len_np`. The current branch is not yet validated to run against that revision. - Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions during `FULL_DECODE_ONLY` replay, speculative-decoding numerical correctness, and EP/TP combinations require model-level validation on the final paired sources. ### How was this patch tested? The latest rebase onto `ad86348b0` resolves an append-location conflict in `test_pcp_manager_v2.py`, retaining both the main DCP padding regression and this PR's tests. AST comparison confirms that all main tests and the final PR tests are preserved, and the other changed files match the automatic merge. Ruff lint/format, syntax parsing, and `git diff --check` passed. Runtime tests were not rerun for this conflict-only update. For the preceding rebase onto `f8481287a`, **6 targeted candidate-method cases passed** in an isolated process in the Ascend container: five main2main draft-sync cases and one mocked v0.28.0 argument-routing case. The exact candidate `propose` method was loaded into the installed runtime classes; native dependencies and the rest of the runtime remained at their installed versions. This is method-level regression evidence, not full rebased-source or dual-version runtime validation. All eight changed Python files passed Ruff lint/format checks, syntax parsing, and `git diff --check`. Full unit-suite and model validation on the final paired sources remain pending. The runtime adaptation suite previously passed on the earlier baseline in an isolated Ascend-container test directory, using CPU tensors and mocks for device kernels: ```bash python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py ``` - **75 passed** (14 existing `torch.jit` deprecation warnings). This is prior-baseline evidence, not a test result for vLLM `7c2f1ff`. - Covers missing/stale dummy PCP state, rank-local mappings, persistent buffer contents and storage, DSA metadata forwarding, non-PCP fallback, graph restrictions, and replicated draft DP synchronization. - The removed validator bypass's dedicated tests are also removed; the existing unrelated platform test is retained. - For the removal, targeted Ruff lint/format checks, Python syntax checks and `git diff --check` passed. No runtime tests were rerun against the upgraded vLLM. The local `format.sh ci` entry point remains unavailable because its shell lacks `pre-commit`. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wzx0726 <zhexuanwu12@gmail.com> Signed-off-by: like-0517 <ithwlike@126.com>
vllm-project#54523 moved the ParallelConfig check into CudaPlatform/RocmPlatform check_and_update_config after this series branched; the series removes the ParallelConfig copy, so remove the platform copies too. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
### What this PR does / why we need it?
#### Change Summary
The PR advances the main2main lane to vLLM v0.30.0 (commit
`ced6857afa0ea7b2e3f0846a62e1394e90f15607`), adapting vllm-ascend to
every upstream change in the `84030bbe` -> `4991f97` -> `ced6857` range.
Because both CI lanes now install vLLM v0.30.0, all
`vllm_version_is("0.29.0")` forks are permanently false and are
collapsed to the main behavior.
| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
| `.github/vllm-main-verified.commit` | — | Updated verified main commit
hash `84030bbe` -> `4991f97` -> `ced6857` (v0.30.0 tag commit) |
| `.github/vllm-release-tag.commit` | — | Bumped the release boundary to
`v0.30.0` |
| `.github/workflows/pr_test.yaml` | — | Added a dual-version cpu-ut
matrix (`vllm_versions`) for the main2main lane; temporarily commented
out pre-commit/mypy and set `cpu-ut` to `if: false`; dropped the
`needs.cpu-ut` requirement from the ready gate; added a
main2main-specific cpu-ut failure hint |
| `Dockerfile` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.310p` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.310p.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a3` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a3.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a5` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a5.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `README.md` | — | CI notes updated to v0.30.0 |
| `README.zh.md` | — | CI notes updated to v0.30.0 |
|
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_num_nans.py`
| — | `VLLM_VERSION` default 0.29.0 -> 0.30.0 |
| `tests/ut/_310p/test_model_runner_v2_310p.py` | v0.30.0 boundary |
Dropped the `vllm_version_is` gate test; `_needs_kv_cache_zeroing_310p`
always uses `spec_config.use_eagle_block_drop()` |
| `tests/ut/core/test_dyntra_lb_scheduler.py` | — | Removed the 0.29
`KVConnectorBlockState.block_ids` assertions; assert the `req_ids` form
only |
| `tests/ut/core/test_scheduler_connector_block_state.py` | — | Removed
the 0.29 `block_ids` snapshot branch |
| `tests/ut/kv_offload/test_native_cpu_offload.py` | — | Block -> chunk
API: assert `spec.num_chunks` |
| `tests/ut/kv_offload/test_npu_offload_spec.py` | — |
`num_blocks`/`kv_bytes_per_block` -> `num_chunks`/`kv_bytes_per_chunk` |
| `tests/ut/models/test_deepseek_v41_registration.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip; registration is always active |
| `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py` | —
| Removed the 0.29 gate for `mamba_fine_grained_prefix_cache` |
| `tests/ut/patch/platform/test_patch_speculative_config_dspark.py` |
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) | Removed
the DeepSeek V4.1 0.29 `pytest.skip` |
| `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` |
[vllm#54523](vllm-project/vllm#54523) | Deleted
the 0.28.0 PCP+DP validator workaround tests; reworked the
version-routing fixtures |
| `tests/ut/patch/worker/test_patch_dspark_pp.py` | — | Parametrize
`legacy: bool` instead of version strings |
| `tests/ut/quantization/configs/test_modelslim_config.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip |
| `tests/ut/spec_decode/test_dspark_proposer.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip marker |
| `tests/ut/spec_decode/test_eagle_proposer.py` | — | `uses_xdrope_dim`
-> `mrope_num_dims` |
| `tests/ut/test_compressed_prefix_cache.py` | — | `replay_boundaries`
now unconditional |
| `tests/ut/worker/test_encoder_acl_graph.py` | — | `axis_keys=()` now
unconditional |
| `tests/ut/worker/test_model_runner_v1.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skips and the `vllm_version_is` patches |
| `tests/ut/worker/test_model_runner_v2.py` | v0.30.0 boundary |
Version-routing fixture rework |
| `tests/ut/worker/test_pcp_manager_v2.py` |
[vllm#56107](vllm-project/vllm#56107) /
[vllm#53867](vllm-project/vllm#53867) | Dropped
the 0.29 `req_states` coverage / added `padded_num_reqs` coverage;
removed the 0.28/0.29 branches |
| `tests/ut/worker/v2/test_pp_utils.py` | — | Replaced the
version-routing matrix with `use_legacy_spec_pp() is False` |
| `vllm_ascend/_310p/model_runner_310p.py` | — | Removed the 0.29
xdrope-position branch |
| `vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py` | — | Removed
the 0.29 xdrope `target_positions[0]` squeeze |
| `vllm_ascend/_310p/worker/v2/model_runner.py` |
[vllm#57270](vllm-project/vllm#57270) | Removed
the 0.29 `max_seq_len_np` kwarg and the 0.29 eagle-block-drop branch |
| `vllm_ascend/_310p/worker/v2/rope.py` | — | mrope `num_dims =
model_config.mrope_num_dims` (dropped the 0.29 constant 3) |
| `vllm_ascend/_310p/worker/v2/states.py` |
[vllm#56908](vllm-project/vllm#56908) |
`UvaBuffer.uva` property -> method; final v0.30.0 form |
| `vllm_ascend/attention/mla_v1.py` |
[vllm#56181](vllm-project/vllm#56181) | Draft
TND_NTD layout forced via `_EXTRA_CTX.is_draft_model`; dropped the 0.29
gate |
| `vllm_ascend/attention/utils.py` |
[vllm#55353](vllm-project/vllm#55353) /
[vllm#56157](vllm-project/vllm#56157) |
Ascend-owned
`_seq_lens_cpu`/`_num_computed_tokens_cpu`/`dcp_local_seq_lens_cpu` are
now unconditional |
| `vllm_ascend/batch_invariant.py` | — | `reduce_sum` accepts
NumPy-style `axis` + `dtype`, rejects `dim`+`axis` together, forwards
`dtype` to the native fallback |
| `vllm_ascend/core/dyntra_lb_scheduler.py` | — |
`KVConnectorBlockState` always uses `req_ids`+`resolve_block_ids`
(dropped the 0.29 `block_ids` snapshot) |
| `vllm_ascend/core/kv_cache_interface.py` |
[vllm#53906](vllm-project/vllm#53906) | MLA
`get_storage_block_size` override unconditional; dropped the 0.29
`storage_block_size` property |
| `vllm_ascend/core/recompute_scheduler.py` | — | Dropped the xdrope
kwarg and the 0.29 `block_ids` snapshot |
| `vllm_ascend/core/scheduler_profiling_chunk.py` | — | Dropped the 0.29
`block_ids` snapshot |
| `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`
| — | Block -> chunk API (`num_chunks`, `kv_bytes_per_chunk`)
unconditional |
| `vllm_ascend/lora/punica_npu.py` |
[vllm#53555](vllm-project/vllm#53555) + boundary
| `add_lora_logits` per-adapter matmul fallback for heads smaller than
the rank; final state drops the `apply_lora_full_linear` binding (both
supported targets predate #53555) |
| `vllm_ascend/models/__init__.py` |
[vllm#56741](vllm-project/vllm#56741) | V4.1
registration (`DeepseekV41ForCausalLM`/`DSparkModel`) unconditional |
| `vllm_ascend/models/deepseek_v41/engram/embedding.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/engram/hash_state.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/engram/parallel.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/model.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/vl_model.py` |
[vllm#56741](vllm-project/vllm#56741) /
[vllm#56554](vllm-project/vllm#56554) |
`deepseek_v41` imports; drop `IMAGE_PAD_ID`/alignment-pad handling on
main |
| `vllm_ascend/ops/mla.py` |
[vllm#56157](vllm-project/vllm#56157) |
`MLAAttention.supports_pcp_dcp = True` set on the class, unconditional |
| `vllm_ascend/ops/rotary_embedding.py` |
[vllm#56446](vllm-project/vllm#56446) | YaRN
mscale signature adaptation; v0.29 branch collapsed |
| `vllm_ascend/patch/__init__.py` |
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) | Registry
entry for the new draft-EP patch; dropped version gates |
| `vllm_ascend/patch/platform/__init__.py` |
[vllm#56741](vllm-project/vllm#56741 Engram
| `patch_engram_config` imported unconditionally |
| `vllm_ascend/patch/platform/patch_balance_schedule.py` | — | Dropped
the 0.29 `block_ids` snapshot |
| `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` |
[vllm#54736](vllm-project/vllm#54736) + boundary
| Accept/forward `allow_partial_hash_hits`; collapsed the 0.29 gate |
| `vllm_ascend/patch/platform/patch_parallel_config.py` |
[vllm#54523](vllm-project/vllm#54523) | Removed
the 0.29 `_validate_parallel_config` PCP+DP workaround; keeps
`use_sequence_parallel_moe` |
| `vllm_ascend/patch/platform/patch_speculative_config.py` |
[vllm#55914](vllm-project/vllm#55914) without
[vllm#56930](vllm-project/vllm#56930) | Skip
`_verify_with_expert_parallelism` for non-MoE draft
(`runner_type=="draft"`) |
| `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` |
[vllm#54523](vllm-project/vllm#54523) | Removed
the 0.28.0 PCP+DP validation workaround |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py` |
[vllm#53781](vllm-project/vllm#53781) | Ascend
`bind_kv_cache_to_layers` (assign the raw allocation); collapsed the
0.29 gate |
| `vllm_ascend/patch/worker/patch_deepseek_v2.py` |
[vllm#53781](vllm-project/vllm#53781) |
Accept/ignore `index_group_builder`; `SparseMLAIndexGroupBuilder` import
collapse |
| `vllm_ascend/patch/worker/patch_mamba_utils.py` |
[vllm#56898](vllm-project/vllm#56898) |
`GPUInputBatch` import source version-gated, then collapsed to
`gpu_input_batch.InputBatch` |
| `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py` |
[vllm#53781](vllm-project/vllm#53781) | Register
Ascend `bind_kv_cache_to_layers`; expose tuple element 0 for the device
filter; collapsed the 0.29 gate |
| `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | — | Legacy
Spec+PP bypass comment (inactive on v0.30.0) |
| `vllm_ascend/patch/worker/patch_v2/patch_spec_pp.py` |
[vllm#56888](vllm-project/vllm#56888) | Alias
`async_tensor_h2d as async_copy_to_gpu`; collapsed the 0.29 gate |
| `vllm_ascend/patch/worker/patch_v2/patch_uva.py` |
[vllm#56908](vllm-project/vllm#56908) | `uva`
property vs method; final v0.30.0 method form |
| `vllm_ascend/spec_decode/llm_base_proposer.py` |
[vllm#56254](vllm-project/vllm#56254) +
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) + boundary
| Gate the V4.1 DSpark import by `HAS_TRITON`; static
`_draft_embed_accepts_mm` check instead of the runtime `embed_input_ids`
probe; collapsed version gates |
| `vllm_ascend/utils.py` | — | `vllm_version_is` docstring 0.29 -> 0.30;
Kimi MLA custom-op registration unconditional |
| `vllm_ascend/worker/model_runner_v1.py` |
[vllm#56741](vllm-project/vllm#56741) | V4.1 dsa
metadata/cache imports unconditional; removed xdrope position handling |
| `vllm_ascend/worker/v2/aclgraph_utils.py` |
[vllm#51700](vllm-project/vllm#51700) |
`ModelAclGraphManager.__init__` accepts/forwards `ubatch_runner`;
`UBatchRunner` import collapse |
| `vllm_ascend/worker/v2/model_runner.py` |
[vllm#56888](vllm-project/vllm#56888) /
[vllm#53867](vllm-project/vllm#53867) /
[vllm#51700](vllm-project/vllm#51700) /
[vllm#57270](vllm-project/vllm#57270) +
[vllm#56181](vllm-project/vllm#56181) + boundary
| Async-copy alias; pass `BatchExecutionDescriptor` to
`maybe_partition_pcp_batch`; `ubatch_runner`; `make_dummy(is_padding)`;
keep the replicated PCP draft on the global batch; collapsed version
gates. Also restores the `_check_oproj_tp_graph_step` guard and the PD
decode-recompute `gather_batch_req_state` override accidentally removed
by the v0.28.0-boundary cleanup (`e30adae80`) |
| `vllm_ascend/worker/v2/pcp_manager.py` |
[vllm#56888](vllm-project/vllm#56888) /
[vllm#56107](vllm-project/vllm#56107) /
[vllm#53867](vllm-project/vllm#53867) +
[vllm#56181](vllm-project/vllm#56181) + boundary
| Async alias; drop `req_states`; forward `padded_num_reqs`;
`prepare_draft_prefill` no-op / `restore_for_sampling` skip; collapsed
version gates |
| `vllm_ascend/worker/v2/pp_utils.py` | — | `use_legacy_spec_pp()`
returns `False` |
| `vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Route
through `_build_uniform_attn_metadata`/`_build_attn_metadata`; re-hook
the Ascend rotary-positions injection onto the new methods; dropped the
0.29 gate |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Same
`BatchExecutionDescriptor` routing; dropped the 0.29 gate |
| `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Same
routing; dropped the 0.29 gate |
### Does this PR introduce _any_ user-facing change?
### How was this patch tested?
- vLLM main:
vllm-project/vllm@84030bb
---------
Signed-off-by: hfadzxy <starmoon_zhang@163.com>
Co-authored-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
### What this PR does / why we need it?
#### Change Summary
The PR advances the main2main lane to vLLM v0.30.0 (commit
`ced6857afa0ea7b2e3f0846a62e1394e90f15607`), adapting vllm-ascend to
every upstream change in the `84030bbe` -> `4991f97` -> `ced6857` range.
Because both CI lanes now install vLLM v0.30.0, all
`vllm_version_is("0.29.0")` forks are permanently false and are
collapsed to the main behavior.
| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
| `.github/vllm-main-verified.commit` | — | Updated verified main commit
hash `84030bbe` -> `4991f97` -> `ced6857` (v0.30.0 tag commit) |
| `.github/vllm-release-tag.commit` | — | Bumped the release boundary to
`v0.30.0` |
| `.github/workflows/pr_test.yaml` | — | Added a dual-version cpu-ut
matrix (`vllm_versions`) for the main2main lane; temporarily commented
out pre-commit/mypy and set `cpu-ut` to `if: false`; dropped the
`needs.cpu-ut` requirement from the ready gate; added a
main2main-specific cpu-ut failure hint |
| `Dockerfile` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.310p` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.310p.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a3` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a3.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a5` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a5.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `README.md` | — | CI notes updated to v0.30.0 |
| `README.zh.md` | — | CI notes updated to v0.30.0 |
|
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_num_nans.py`
| — | `VLLM_VERSION` default 0.29.0 -> 0.30.0 |
| `tests/ut/_310p/test_model_runner_v2_310p.py` | v0.30.0 boundary |
Dropped the `vllm_version_is` gate test; `_needs_kv_cache_zeroing_310p`
always uses `spec_config.use_eagle_block_drop()` |
| `tests/ut/core/test_dyntra_lb_scheduler.py` | — | Removed the 0.29
`KVConnectorBlockState.block_ids` assertions; assert the `req_ids` form
only |
| `tests/ut/core/test_scheduler_connector_block_state.py` | — | Removed
the 0.29 `block_ids` snapshot branch |
| `tests/ut/kv_offload/test_native_cpu_offload.py` | — | Block -> chunk
API: assert `spec.num_chunks` |
| `tests/ut/kv_offload/test_npu_offload_spec.py` | — |
`num_blocks`/`kv_bytes_per_block` -> `num_chunks`/`kv_bytes_per_chunk` |
| `tests/ut/models/test_deepseek_v41_registration.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip; registration is always active |
| `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py` | —
| Removed the 0.29 gate for `mamba_fine_grained_prefix_cache` |
| `tests/ut/patch/platform/test_patch_speculative_config_dspark.py` |
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) | Removed
the DeepSeek V4.1 0.29 `pytest.skip` |
| `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` |
[vllm#54523](vllm-project/vllm#54523) | Deleted
the 0.28.0 PCP+DP validator workaround tests; reworked the
version-routing fixtures |
| `tests/ut/patch/worker/test_patch_dspark_pp.py` | — | Parametrize
`legacy: bool` instead of version strings |
| `tests/ut/quantization/configs/test_modelslim_config.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip |
| `tests/ut/spec_decode/test_dspark_proposer.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip marker |
| `tests/ut/spec_decode/test_eagle_proposer.py` | — | `uses_xdrope_dim`
-> `mrope_num_dims` |
| `tests/ut/test_compressed_prefix_cache.py` | — | `replay_boundaries`
now unconditional |
| `tests/ut/worker/test_encoder_acl_graph.py` | — | `axis_keys=()` now
unconditional |
| `tests/ut/worker/test_model_runner_v1.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skips and the `vllm_version_is` patches |
| `tests/ut/worker/test_model_runner_v2.py` | v0.30.0 boundary |
Version-routing fixture rework |
| `tests/ut/worker/test_pcp_manager_v2.py` |
[vllm#56107](vllm-project/vllm#56107) /
[vllm#53867](vllm-project/vllm#53867) | Dropped
the 0.29 `req_states` coverage / added `padded_num_reqs` coverage;
removed the 0.28/0.29 branches |
| `tests/ut/worker/v2/test_pp_utils.py` | — | Replaced the
version-routing matrix with `use_legacy_spec_pp() is False` |
| `vllm_ascend/_310p/model_runner_310p.py` | — | Removed the 0.29
xdrope-position branch |
| `vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py` | — | Removed
the 0.29 xdrope `target_positions[0]` squeeze |
| `vllm_ascend/_310p/worker/v2/model_runner.py` |
[vllm#57270](vllm-project/vllm#57270) | Removed
the 0.29 `max_seq_len_np` kwarg and the 0.29 eagle-block-drop branch |
| `vllm_ascend/_310p/worker/v2/rope.py` | — | mrope `num_dims =
model_config.mrope_num_dims` (dropped the 0.29 constant 3) |
| `vllm_ascend/_310p/worker/v2/states.py` |
[vllm#56908](vllm-project/vllm#56908) |
`UvaBuffer.uva` property -> method; final v0.30.0 form |
| `vllm_ascend/attention/mla_v1.py` |
[vllm#56181](vllm-project/vllm#56181) | Draft
TND_NTD layout forced via `_EXTRA_CTX.is_draft_model`; dropped the 0.29
gate |
| `vllm_ascend/attention/utils.py` |
[vllm#55353](vllm-project/vllm#55353) /
[vllm#56157](vllm-project/vllm#56157) |
Ascend-owned
`_seq_lens_cpu`/`_num_computed_tokens_cpu`/`dcp_local_seq_lens_cpu` are
now unconditional |
| `vllm_ascend/batch_invariant.py` | — | `reduce_sum` accepts
NumPy-style `axis` + `dtype`, rejects `dim`+`axis` together, forwards
`dtype` to the native fallback |
| `vllm_ascend/core/dyntra_lb_scheduler.py` | — |
`KVConnectorBlockState` always uses `req_ids`+`resolve_block_ids`
(dropped the 0.29 `block_ids` snapshot) |
| `vllm_ascend/core/kv_cache_interface.py` |
[vllm#53906](vllm-project/vllm#53906) | MLA
`get_storage_block_size` override unconditional; dropped the 0.29
`storage_block_size` property |
| `vllm_ascend/core/recompute_scheduler.py` | — | Dropped the xdrope
kwarg and the 0.29 `block_ids` snapshot |
| `vllm_ascend/core/scheduler_profiling_chunk.py` | — | Dropped the 0.29
`block_ids` snapshot |
| `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`
| — | Block -> chunk API (`num_chunks`, `kv_bytes_per_chunk`)
unconditional |
| `vllm_ascend/lora/punica_npu.py` |
[vllm#53555](vllm-project/vllm#53555) + boundary
| `add_lora_logits` per-adapter matmul fallback for heads smaller than
the rank; final state drops the `apply_lora_full_linear` binding (both
supported targets predate #53555) |
| `vllm_ascend/models/__init__.py` |
[vllm#56741](vllm-project/vllm#56741) | V4.1
registration (`DeepseekV41ForCausalLM`/`DSparkModel`) unconditional |
| `vllm_ascend/models/deepseek_v41/engram/embedding.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/engram/hash_state.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/engram/parallel.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/model.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/vl_model.py` |
[vllm#56741](vllm-project/vllm#56741) /
[vllm#56554](vllm-project/vllm#56554) |
`deepseek_v41` imports; drop `IMAGE_PAD_ID`/alignment-pad handling on
main |
| `vllm_ascend/ops/mla.py` |
[vllm#56157](vllm-project/vllm#56157) |
`MLAAttention.supports_pcp_dcp = True` set on the class, unconditional |
| `vllm_ascend/ops/rotary_embedding.py` |
[vllm#56446](vllm-project/vllm#56446) | YaRN
mscale signature adaptation; v0.29 branch collapsed |
| `vllm_ascend/patch/__init__.py` |
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) | Registry
entry for the new draft-EP patch; dropped version gates |
| `vllm_ascend/patch/platform/__init__.py` |
[vllm#56741](vllm-project/vllm#56741 Engram
| `patch_engram_config` imported unconditionally |
| `vllm_ascend/patch/platform/patch_balance_schedule.py` | — | Dropped
the 0.29 `block_ids` snapshot |
| `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` |
[vllm#54736](vllm-project/vllm#54736) + boundary
| Accept/forward `allow_partial_hash_hits`; collapsed the 0.29 gate |
| `vllm_ascend/patch/platform/patch_parallel_config.py` |
[vllm#54523](vllm-project/vllm#54523) | Removed
the 0.29 `_validate_parallel_config` PCP+DP workaround; keeps
`use_sequence_parallel_moe` |
| `vllm_ascend/patch/platform/patch_speculative_config.py` |
[vllm#55914](vllm-project/vllm#55914) without
[vllm#56930](vllm-project/vllm#56930) | Skip
`_verify_with_expert_parallelism` for non-MoE draft
(`runner_type=="draft"`) |
| `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` |
[vllm#54523](vllm-project/vllm#54523) | Removed
the 0.28.0 PCP+DP validation workaround |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py` |
[vllm#53781](vllm-project/vllm#53781) | Ascend
`bind_kv_cache_to_layers` (assign the raw allocation); collapsed the
0.29 gate |
| `vllm_ascend/patch/worker/patch_deepseek_v2.py` |
[vllm#53781](vllm-project/vllm#53781) |
Accept/ignore `index_group_builder`; `SparseMLAIndexGroupBuilder` import
collapse |
| `vllm_ascend/patch/worker/patch_mamba_utils.py` |
[vllm#56898](vllm-project/vllm#56898) |
`GPUInputBatch` import source version-gated, then collapsed to
`gpu_input_batch.InputBatch` |
| `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py` |
[vllm#53781](vllm-project/vllm#53781) | Register
Ascend `bind_kv_cache_to_layers`; expose tuple element 0 for the device
filter; collapsed the 0.29 gate |
| `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | — | Legacy
Spec+PP bypass comment (inactive on v0.30.0) |
| `vllm_ascend/patch/worker/patch_v2/patch_spec_pp.py` |
[vllm#56888](vllm-project/vllm#56888) | Alias
`async_tensor_h2d as async_copy_to_gpu`; collapsed the 0.29 gate |
| `vllm_ascend/patch/worker/patch_v2/patch_uva.py` |
[vllm#56908](vllm-project/vllm#56908) | `uva`
property vs method; final v0.30.0 method form |
| `vllm_ascend/spec_decode/llm_base_proposer.py` |
[vllm#56254](vllm-project/vllm#56254) +
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) + boundary
| Gate the V4.1 DSpark import by `HAS_TRITON`; static
`_draft_embed_accepts_mm` check instead of the runtime `embed_input_ids`
probe; collapsed version gates |
| `vllm_ascend/utils.py` | — | `vllm_version_is` docstring 0.29 -> 0.30;
Kimi MLA custom-op registration unconditional |
| `vllm_ascend/worker/model_runner_v1.py` |
[vllm#56741](vllm-project/vllm#56741) | V4.1 dsa
metadata/cache imports unconditional; removed xdrope position handling |
| `vllm_ascend/worker/v2/aclgraph_utils.py` |
[vllm#51700](vllm-project/vllm#51700) |
`ModelAclGraphManager.__init__` accepts/forwards `ubatch_runner`;
`UBatchRunner` import collapse |
| `vllm_ascend/worker/v2/model_runner.py` |
[vllm#56888](vllm-project/vllm#56888) /
[vllm#53867](vllm-project/vllm#53867) /
[vllm#51700](vllm-project/vllm#51700) /
[vllm#57270](vllm-project/vllm#57270) +
[vllm#56181](vllm-project/vllm#56181) + boundary
| Async-copy alias; pass `BatchExecutionDescriptor` to
`maybe_partition_pcp_batch`; `ubatch_runner`; `make_dummy(is_padding)`;
keep the replicated PCP draft on the global batch; collapsed version
gates. Also restores the `_check_oproj_tp_graph_step` guard and the PD
decode-recompute `gather_batch_req_state` override accidentally removed
by the v0.28.0-boundary cleanup (`e30adae80`) |
| `vllm_ascend/worker/v2/pcp_manager.py` |
[vllm#56888](vllm-project/vllm#56888) /
[vllm#56107](vllm-project/vllm#56107) /
[vllm#53867](vllm-project/vllm#53867) +
[vllm#56181](vllm-project/vllm#56181) + boundary
| Async alias; drop `req_states`; forward `padded_num_reqs`;
`prepare_draft_prefill` no-op / `restore_for_sampling` skip; collapsed
version gates |
| `vllm_ascend/worker/v2/pp_utils.py` | — | `use_legacy_spec_pp()`
returns `False` |
| `vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Route
through `_build_uniform_attn_metadata`/`_build_attn_metadata`; re-hook
the Ascend rotary-positions injection onto the new methods; dropped the
0.29 gate |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Same
`BatchExecutionDescriptor` routing; dropped the 0.29 gate |
| `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Same
routing; dropped the 0.29 gate |
### Does this PR introduce _any_ user-facing change?
### How was this patch tested?
- vLLM main:
vllm-project/vllm@84030bb
---------
Signed-off-by: hfadzxy <starmoon_zhang@163.com>
Co-authored-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
Purpose
Move the current PCP-with-DP capability check from the platform-independent ParallelConfig validator to the GPU MRV2 PCPManager. This keeps the existing GPU behavior while allowing out-of-tree hardware backends to provide their own PCP+DP support.
Test Plan
Test Result
Essential Elements of an Effective PR Description Checklist